Papers with corpus similarity measures
Predicting Embedding Reliability in Low-Resource Settings Using Corpus Similarity Measures (2022.lrec-1)
Copied to clipboard
| Challenge: | a paper aims to evaluate embedding similarity, stability and reliability in low-resource settings . it uses corpus similarity measures before training to predict properties of embeddables . |
| Approach: | They use corpus similarity measures before training to predict properties of embeddings . they then apply the same measures to low-resource settings by modelling reliability . authors hope to use this method to evaluate low-source languages with limited corpus size . |
| Outcome: | The paper shows that it is possible to predict downstream embedding similarity using upstream corpus similarity measures . the main finding is that the measures remain robust on small amounts of training data . |
Validating and Exploring Large Geographic Corpora (2024.lrec-main)
Copied to clipboard
| Challenge: | a paper examines the impact of corpus creation decisions on multi-lingual web corpora . the goal is to understand the impact on downstream corporata with a focus on under-represented languages and populations. |
| Approach: | This paper evaluates the impact of corpus creation decisions on multi-lingual web corpora . three cleaning methods are used to improve the quality of sub-corpora in the common crawl . the goal is to understand the impact on downstream corporan with a focus on under-represented languages . |
| Outcome: | The results show that the validity of sub-corpora is improved with each stage of cleaning but that this improvement is unevenly distributed across languages and populations. |